Skip to content

Add NSECE childcare attendance candidate and validation tools - #916

Draft
hua7450 wants to merge 20 commits into
PolicyEngine:mainfrom
hua7450:fix/us-childcare-attendance
Draft

hua7450 wants to merge 20 commits into
PolicyEngine:mainfrom
hua7450:fix/us-childcare-attendance

Conversation

@hua7450

@hua7450 hua7450 commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

Missing person-level childcare attendance can leave household CCDF calculations at zero even when childcare expenses are present. This PR adds a dataset-side NSECE attendance adapter and fiscal-build integration, preserving PolicyEngine-US input defaults.

Draft — not ready to merge. Maria's September 22 review closes the prior integrity, row-completeness, household-identity and provenance findings. The raw-source build/certification requirement (C2) and household-model validation concern (A3) remain open. No population is published.

Implementation

  • Verify the two pinned 2024 NSECE household/calendar files. Derive joint days and hours from classified 15-minute care blocks, preserve measured nonattendance, exclude K–8 schooling, and retain ambiguous/incomplete calendars as unknown.
  • Preserve noncalendar respondents' observed regular hours while modeling days and irregular care. Transfer schedules jointly to children ages 0–12 using age, parent work, region and income; preserve observed cells, stable source identities, clones, existing population values and weights.
  • Run attendance after the childcare-expense and hours producers. Require all three export columns and bind exact person identities, ages, membership and attendance to source/code/runtime/settings receipts. Native loaders, fiscal exports, L0 exports and annual projections preserve and check that binding.
  • Accept pinned source configuration through the exact-k launcher. The runbook explicitly requires this configuration for multispine ingress, which does not restore an attendance receipt.

September 27 review fixes

Merged PolicyEngine main at 9e5b0cee, resolving the builder, generated coverage and test-layout conflicts. Tests now use the repository's engine-free/US-engine directories and shared test support.

A4: revalidate attendance immediately before target materialization and include its binding digest in the checkpoint identity. Reform vectors already include the full materializer identity, so they also invalidate. Both changed values and a changed recipe that emits identical values miss the old caches; unchanged executions reuse them. A negative-control run reproduces the stale recipe-only checkpoint hit without this fix.

Preserve the attendance receipt before main's new post-export scorer hashes the completed H5, while retaining its stored-input checks and batching behavior.

Evidence and remaining work

The September 17 candidate contains 166,321 people, 57,240 households and 31,889 under-13 children. Fresh September 27 verification through both native loaders reproduces the committed integrity report byte for byte, preserving every original value/weight and the exact attendance binding. The source recipe and population estimates are unchanged.

The earlier attendance-only 2026-policy comparison reduced zero-benefit jurisdictions from 31 to 2 (MD/NV), with potential modeled benefits rising from $2.254B to $6.206B on fixed source ages/incomes. These are diagnostic estimates, not calibrated spending, caseload estimates or model-accuracy evidence.

Current review audit and validation · Aggregate reports and experiment history

Validation

  • Affected suite: 1,448 passed, four fixture failures and two pinned-feed skips. After fixing the fixtures, all 23 serializer tests and 17 selected fiscal-builder/cache tests pass, covering every failure. The entire broad suite was not repeated after those test-only fixes.
  • Focused attendance suite: 145 passed with the US engine; 138 passed separately without any country engine installed. Both cache regressions also pass in that engine-free environment.
  • All six wheels build and pass content inspection. Compiler coverage is 42,187/42,187 fields with 41/41 inventory checks. Lint, changed-file formatting, whitespace and the 540-module CI inventory pass.
  • Existing native candidate integrity report reproduced byte for byte. This does not replace the outstanding raw-source build or population validation.

Full details are in the linked audit. GitHub CI has not been monitored for this update.

Survey data for reviewers

Download the source ZIP from Google Drive, accessible to signed-in PolicyEngine Google accounts. NSECE-2024-PR916-source-files.zip contains:

  • 39466-0004-Data.tsv (DS0004): childcare calendar records.
  • 39466-0005-Data.tsv (DS0005): household/child characteristics, survey weights and regular childcare hours.
  • README-FILES-USED.txt: roles, SHA-256 checksums and this PR link.

These are the unchanged source bytes used by the adapter. ASEC files, target populations and private person-level receipt inventories are not included. Only synthetic tests and aggregate evidence are committed.

Addresses #915; does not close the remaining statistical, provider/activity or older-child gaps.

@hua7450 hua7450 changed the title Add child-level childcare attendance donor preparation Add NSECE child attendance adapter and candidate validation Sep 12, 2026
@hua7450 hua7450 changed the title Add NSECE child attendance adapter and candidate validation Integrate NSECE child attendance and validate full US population Sep 13, 2026
@hua7450
hua7450 marked this pull request as ready for review September 14, 2026 17:14
@hua7450

hua7450 commented Sep 14, 2026

Copy link
Copy Markdown
Contributor Author

Hi @juaristi22, could you please review this PR when you have a chance? This is my first time working in a data repo, and I’m not yet confident that my approach to adding the childcare attendance inputs follows Microcosm’s requirements and conventions.

Please be as critical and thorough as you would normally be—don’t hold back because it’s my first contribution here. I’d especially appreciate feedback on the survey mapping, imputation approach, build integration, and whether the validation is sufficient. Please also point out even minor issues with wording, naming, formatting, or code organization; I want to learn the repo’s standards and get this right.

The PR description includes a “Survey data for reviewers” section with the two source files and a README explaining how they are used. The Drive folder is accessible with a signed-in PolicyEngine Google account. The remaining limitations and the separate integration needed for #893 are documented as well.

Thank you!

@juaristi22

juaristi22 commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Program review

Base repository: PolicyEngine/microcosm
PR number: 916
Reviewed head SHA: 3fe3e68
Merge base SHA: a9cc63e
Mode: full
Scope: changed behavior and affected dependencies
Source manifest: /private/tmp/policyengine-command-runs/d374f57aad22/pr-916-review-sources.json
Review status: PARTIAL

Source Documents

No source documents registered; see scope and validation.

What looks good

  • The PR addresses the problem at the correct layer: it supplies measured and modeled attendance through the population build instead of changing PolicyEngine-US defaults or inferring attendance from expenses, eligibility, or benefit receipt.
  • The survey adapter is deliberately fail-closed. It hash-verifies both NSECE inputs, distinguishes missing/partial/ambiguous calendars from measured nonattendance, excludes K–8 schooling, treats regular and irregular care separately, and derives the three attendance fields jointly from the classified calendar.
  • The noncalendar bridge preserves each child's observed regular weekly hours, borrows days and irregular-care intensity jointly, requires compatible participation, retains all ties at the nearest-donor cutoff, uses survey weights, and never feeds reconstructed records back into its donor pool. The masked-calendar assessment holds out entire households, avoiding household leakage.
  • The population transfer follows Microcosm's core data contracts well: typed design weights, exact age in every fallback, joint schedule draws, stable source-person identities, clone reconciliation, observed-cell preservation (including zero), and per-cell donor provenance. The sibling mixture is also a meaningful improvement over independent child draws, even though its intensity validation should be expanded as noted below.
  • ASEC harmonization uses actual within-household parent pointers and last-week work rather than treating every adult or earner as a parent. Region comes from shared Census constants, and restored income values are protected by file hashes plus person/year/age/line reconciliation while original raw columns remain unchanged.
  • Build integration is thoughtfully scoped. The stage runs after the existing childcare-expense producer; all three inputs enter the release-coverage contract; the native exporter writes a new file, reloads it, and verifies the pre-existing entity tables, dtypes, household weights, and time period.
  • The validation narrative is unusually candid. It reports source attrition, fallback usage, household-separated diagnostics, all 51 jurisdiction results, and the remaining California/Maryland/Nevada blockers. It consistently labels the result as an attendance-only counterfactual, retains production_ready: false, and does not claim a calibrated release, measured spending, or a change to the default population.

Critical

C1 — CRITICAL (Must Fix): attendance values are not bound to their source receipt, and refreshing an enriched base silently retains stale values (OPEN)

Location: tools/build_us_fiscal_refresh_release.py:1567-1575,9349-9361,11801-11812; packages/microcosm-build/src/microcosm/build/us_runtime/childcare_attendance.py:164-167,208-210; packages/microcosm-build/src/microcosm/build/us_runtime/release_input_coverage.py:567-571,588-612; packages/microcosm-build/src/microcosm/build/us_runtime/l0_refit_export.py:504-523.

Trigger / reproduction: use an otherwise valid base H5 containing non-default values for the three attendance columns (the native candidate produced by this PR is such an H5) and run the fiscal builder without the two NSECE TSV flags. The argument parser only checks that the flags are paired when one is supplied, the stage is skipped when both are absent, the generic H5 loader reads only the six entity tables and discards _childcare_attendance_receipt, and the release-input gate checks only presence/non-default signal. us_source_coverage.json adds childcare evidence only when optional in-memory metadata happens to exist. The build can therefore accept the attendance columns without authenticating the pinned NSECE hashes, seed, matching recipe, modeled-age boundary, or outside-domain policy. A second trigger is to supply new TSVs or a new seed on a previously enriched frame: the imputer labels every pre-existing non-null cell observed, skips every complete source person, then overwrites frame-level metadata with the new source receipt/seed even though the old values were retained.

Expected: a release must either execute the pinned source stage or validate and carry a receipt cryptographically/content-bound to the exact persisted attendance values. Re-running with a changed source identity or seed must either recompute derived cells or reject the incompatible pre-existing provenance.

Observed: arbitrary or stale non-default values satisfy the hard coverage manifest, while source coverage may contain no attendance receipt; when the stage is requested on an enriched input, output values and new metadata can describe different executions.

Impact: this defeats the repository's load-bearing artifact/provenance contract and can certify materially different state childcare-subsidy outputs as if they came from the reviewed NSECE mapping. The committed comparison shows the attendance inputs move potential modeled benefits from about $2.25B to $5.29B, so accepting unbound/stale values is output-material. The new tests cover flag pairing and same-input idempotence, but not receipt-required loading, changed-source/seed refresh, or final source-coverage enforcement.

Should Address

A1 — SHOULD ADDRESS: the final release gate does not reassert row-complete attendance (OPEN)

  • Location: packages/microcosm-build/src/microcosm/build/us_runtime/childcare_attendance_stage.py:125-127; packages/microcosm-build/src/microcosm/build/us_runtime/release_input_coverage.py:559-612.
  • Trigger: A later build operation, schema conversion, or future refactor introduces null/default attendance in only part of the export population after the NSECE stage has completed.
  • Expected: Because the documentation says unresolved under-13 values fail export and the attendance stage's contract requires all three fields to be complete and internally valid, the final pre-write release boundary should re-run that rowwise contract (or have an equivalent attendance-specific gate).
  • Observed: with_us_childcare_attendance_inputs calls assert_childcare_attendance_exportable at its own stage boundary, but the final generic input-coverage gate only requires each named column to have at least one finite, non-default observation. A partially null attendance column therefore remains nondegenerate and passes this final gate.
  • Impact: The present pipeline is guarded when the stage returns, but the release contract cannot detect partial loss between that point and final serialization. That is a missing boundary check on the precise behavior this PR claims to enforce; a later defect could ship rows that fall back to engine defaults.
  • Evidence: The gate documents and implements its column-level criterion at release_input_coverage.py:565-573 and :588-612; the attendance-specific complete-row assertion is only at childcare_attendance_stage.py:127 (and in the standalone native exporter at :169).

A2 — SHOULD ADDRESS: invalid or missing household source identities collapse into one sibling-dependence group (OPEN)

Location: packages/microcosm-build/src/microcosm/build/us_runtime/childcare_population.py:95-100; packages/microcosm-build/src/microcosm/build/us_runtime/childcare_attendance.py:113-115,232-248.

Trigger / reproduction: pass a US frame whose household_source_id column exists but contains a null, or whose person-to-household mapping fails. Harmonization maps the value and immediately calls .astype(str), converting null to the accepted nonempty ID "nan". With fitted sibling_dependence > 0, every affected household uses the same childcare_household_rank hash.

Expected: source household identities used to couple sibling draws should be complete and every person link should resolve; invalid identity must fail closed, as comparable source-ID mapping code in puf_support.py:795-809 does.

Observed: unrelated children with unresolved source identities are treated as one synthetic household for the shared-rank component.

Impact: on malformed or legacy inputs this introduces artificial cross-household dependence and masks an upstream linkage defect. The qualified BuildP parent likely satisfies the late-producer finite-ID invariant, so this does not refute the committed candidate, but the new public stage itself does not enforce its stated identity precondition.

A3 — SHOULD ADDRESS: sibling validation covers only binary participation, while the shared rank couples full schedule intensity (OPEN)

  • Location: packages/microcosm-build/src/microcosm/build/us_runtime/nsece_childcare_dependence.py:17-29,32-79; packages/microcosm-build/src/microcosm/build/us_runtime/childcare_attendance.py:171-180,232-256; packages/microcosm-build/src/microcosm/build/us_runtime/nsece_childcare_assessment.py:103-133.
  • Trigger: A target household has multiple under-13 siblings and enters the fitted shared-rank branch (rho = 0.7413 in the committed artifact).
  • Expected: Validation should cover every material dependence the mechanism induces—at least joint participation, joint days, and joint weekly hours—and show how households with more than two children behave, because attendance days/hours feed subsidy eligibility and amounts.
  • Observed: The fit deliberately selects only the youngest two children per fully observed household and estimates one rho from whether both have positive days. The implementation then sorts each donor pool by days and hours and may apply the same household quantile to every sibling, coupling schedule intensity as well as participation and extending the two-child estimate to larger sibships. Cross-validation compares only the both in care probability. The committed aggregate is encouraging for that one target (observed 0.3286, independent 0.2585, fully shared 0.3530), but it does not validate the added hours/days dependence (experiments/us-childcare-attendance/qualified-preparation.json:400-407).
  • Impact: Marginal child distributions remain preserved by the quantile construction, but household-level schedule intensity can be too strongly correlated even when the binary joint-attendance moment fits. Nonlinear state eligibility/benefit outcomes can therefore differ without this diagnostic revealing it.
  • Evidence: Code trace above and the report's own limitation that sibling assignments still need validation (qualified-preparation.json:393-398). Add held-out weighted joint moments/correlations for days and weekly hours, plus an explicit diagnostic for 3+ child households, before treating the fitted household process as qualified.

Suggestions

S1 — SUGGESTION: foreground the questionnaire-transport discrepancy and define acceptance criteria (OPEN)

  • Location: experiments/us-childcare-attendance/qualified-preparation.json:1115-1148; docs/us-childcare-attendance.md:55-68,154-169.
  • Trigger: A reviewer reads “qualified” as evidence that the summer/typical-May and new-school-year instruments are exchangeable with the main calendar instrument conditional on the matching cells.
  • Expected: The main narrative should state the observed transport discrepancy numerically and define what level/sensitivity would be acceptable for the candidate's intended use.
  • Observed: The report estimates weighted regular hours of 12.38 versus a calendar-sample conditional expectation of 14.80 for questionnaire 2, and 10.16 versus 13.92 for questionnaire 3 (about 20% and 37% higher conditional expectations, respectively). The code correctly preserves each noncalendar child's observed regular hours, so those differences do not overwrite that measured field; however, they are evidence that instrument/sample transport is not innocuous for the donor-based days and irregular-care assumptions. The limitations are present and production_ready is false, but no acceptance rule or sensitivity bound translates the discrepancy into a qualification decision.
  • Impact: This is not a demonstrated wrong output, but readers can overinterpret the strong masked-calendar mean comparison as validation on the actual noncalendar population. Promote these numbers into the README/docs summary, state that days and irregular care remain unidentified, and report benefit sensitivity to alternative irregular-care/day transport choices.
  • Evidence: The assessment itself expressly says the regular-hours comparison cannot validate days or irregular care (packages/microcosm-build/src/microcosm/build/us_runtime/nsece_childcare_assessment.py:163-211) and that the masked test reconstructs cases whose calendars actually exist rather than the other instruments (:215-292).

S2 — SUGGESTION: the persisted operation order collapses three distinct transformations to one generic label (OPEN)

  • Location: packages/microcosm-build/src/microcosm/build/us/childcare_attendance_source.json:10-14; packages/microcosm-build/src/microcosm/build/us_runtime/childcare_attendance_stage.py:112-146; committed example at experiments/us-childcare-attendance/qualified-population-comparison.json:32-59.
  • Trigger: A release reviewer or diagnostic consumer uses the candidate receipt to reconstruct which source operations ran and in what order.
  • Expected: The receipt should retain the distinct declared operations (calendar_attendance, regular_hours_schedule_bridge, joint_weighted_schedule_transfer) and identify the data-fitted sibling-dependence step.
  • Observed: operation_order serializes operation.kind, and all three manifest records use the same kind, derive_childcare_inputs. The committed receipt consequently contains that label three times. The fitted sibling_dependence object is separately present, so the numeric fit is not lost, but the structured operation chain is incomplete/ambiguous.
  • Impact: This does not change modeled values, but it weakens provenance, makes receipt-level audits harder, and can hide future reordering or insertion of material data-dependent steps.
  • Evidence: Serialize the operation-specific operation parameter (ideally the whole normalized operation record) and include fit_sibling_dependence as an explicit material transform.

Coordinator assessment: The code reviewer independently identified the same receipt-auditability issue; it is consolidated here once.

Evidence Gaps

  • Required local tests were NOT RUN because the snapshot had no existing environment with pytest/pandas and the review contract prohibited installing dependencies; the exact-head GitHub CI evidence is green.
  • The licensed DS4/DS5 TSVs and exact 2024 NSECE value-label codebooks were unavailable, so provider/gap/missingness/respondent-care classifications, file hashes, weight universes, and the NSECE-to-ASEC parent-work crosswalk could not be independently verified.
  • The May/fall noncalendar instruments do not observe days or irregular care; the masked-calendar test validates reconstruction on main-instrument records but cannot establish transport validity for the actual noncalendar population.
  • The licensed NSECE inputs, pinned ASEC cache, parent H5, and candidate artifacts were unavailable, so the real-data attrition, population enrichment, and 51-jurisdiction comparison could not be independently reproduced.

Notes

  • The PR head is 46 commits behind current main; this is informational. Only geography_constants.py changed on both sides, and Git reports the current merge as clean.
  • All 23 GitHub checks reported SUCCESS at the exact reviewed head.
  • The PR deliberately does not publish a replacement population, change PolicyEngine-US defaults, fix state-specific provider/activity inputs, or resolve the separate household-benefit aggregation issue.

Validation Summary

Inspected the five-commit, 36-file merge-base diff and affected runtime, build, serializer, coverage, documentation, tests, and aggregate evidence. Local tests: NOT RUN (environment unavailable; no dependency install). GitHub CI: 23/23 SUCCESS at head 3fe3e68. Official public source review covered the NSECE study page, Census ASEC variable catalog, and BLS CPI-U table; licensed source bytes/codebooks and full-data artifacts were unavailable.

Timing

setup seconds: 29.00s; scope seconds: 67.00s; parallel review seconds: 765.00s; policy role seconds: 765.00s; code role seconds: 525.00s; adjudication seconds: 0.00s; consolidation cleanup seconds: 140.00s; elapsed seconds: 976.00s

Review Severity

REQUEST_CHANGES. Open findings: 1 critical, 3 should address, 2 suggestions.

@hua7450
hua7450 marked this pull request as draft September 15, 2026 18:32
@hua7450 hua7450 changed the title Integrate NSECE child attendance and validate full US population Add NSECE childcare attendance candidate and validation tools Sep 16, 2026
hua7450 and others added 9 commits September 16, 2026 20:05
Resolve the coverage report and multispine test conflicts using the specification generated from the combined branch. Preserve childcare recipe and draft validation limitations.
- Keep frame metadata through the ACA source-output step so the attendance
  receipt reaches the final export check.
- Add the native receipt key without rewriting the person table, and persist
  and restore only the attendance context and binding.
- Compare recipe code/runtime identity at bind and release export only;
  read-only native ingress checks content.
- Carry the receipt through the L0 refit export and restore it in the fiscal
  builder's --base-h5 loader.
- Refuse a build with neither NSECE source files nor bound attendance before
  calibration, and keep an unbound-attendance failure in the batched report.
- Widen thin noncalendar-bridge cells to the next matching level and count
  children by matching level.
- Reject unlisted region, parent-work and negative income codes (User's Guide
  HH-63, HH-175, HH-483).
- Fail the require_observed outside-domain policy early with the fixing flag,
  and run the stage after the hours producer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Rebuild from the pinned parent through the production stage, then rerun the
population comparison, paired sensitivity and native-loader verification.
Add a five-split masked-calendar comparison of the bridge before and after
thin-cell widening. Earlier reports remain as historical evidence.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Classify write_native_childcare_receipt as a non-table HDF write, let the ACA
source-output test's Frame stub carry metadata and assert it survives, and fix
two new tests that indexed the wrong source row and used a frame whose adult
already had observed attendance.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@hua7450

hua7450 commented Sep 21, 2026

Copy link
Copy Markdown
Contributor Author

Hi @juaristi22, thanks again for the detailed review. Could you please take another look at the latest revision (25a54a4e)?

I've implemented fixes for the receipt/value binding (C1), final row-completeness checks (A1), invalid household identities (A2), and explicit operation provenance (S2). Attendance receipts also survive native exports and annual projections. All 24 CI checks pass.

The expanded household and transport diagnostics still show unresolved statistical limitations related to A3/S1. The required raw-source build and certification are also outstanding, with upstream prerequisites documented in the merge-readiness audit. I'm keeping this as a draft.

I'd especially appreciate your assessment of whether the engineering fixes meet the repo's requirements and what evidence or methodological changes are needed to address the remaining validation gaps. Please keep being thorough, including on wording, naming, formatting, and process—I'm still learning the data repo's conventions. Thank you!

@juaristi22

Copy link
Copy Markdown
Collaborator

Program review

Base repository: PolicyEngine/microcosm
PR number: 916
Reviewed head SHA: 25a54a4
Merge base SHA: 18c6d39
Mode: incremental from 3fe3e68
Scope: changed behavior and affected dependencies
Source manifest: /private/tmp/policyengine-command-runs/d374f57aad22/pr-916-review-sources.json
Review status: PARTIAL

Source Documents

No source documents registered; see scope and validation.

Critical

C1 — CRITICAL — RESOLVED: attendance values are content-bound to their execution and checked through output boundaries (RESOLVED)

  • Location: packages/microcosm-build/src/microcosm/build/us_runtime/childcare_attendance_receipt.py:47-72,85-145,148-203,206-250; packages/microcosm-build/src/microcosm/build/us_runtime/childcare_attendance_stage.py:128-145,229-326; packages/microcosm-build/src/microcosm/build/us_runtime/h5_io.py:930-962; packages/microcosm-build/src/microcosm/build/us_runtime/l0_refit_export.py:490-529; packages/microcosm-build/src/microcosm/build/us_annual_static_aging.py:186-221,275-298,390-391; tools/build_us_fiscal_refresh_release.py:2662-2679,11712-11728.
  • Prior trigger: a supplied enriched H5 could retain arbitrary or stale attendance while losing or replacing its receipt; a changed source/seed could be described by new metadata without recomputing the old values.
  • Expected / observed now: each receipt binds source pins, the production contract, recipe-module hashes, runtime versions, stage settings, and execution context to every retained person's ID, household membership, age, and three attendance values. Existing source metadata must have a valid binding; changed source/settings are rejected; nonzero native attendance without a receipt is rejected. Both native loaders restore and validate the receipt, the fiscal build refuses before calibration and immediately before write, and the written H5 is reloaded before evidence is emitted. Annual projection restores the base receipt, writes it to each projected H5, and checks the retained identities/ages/attendance after serialization while allowing weights to change.
  • Evidence / impact: experiments/us-childcare-attendance/review-fixes-verification.json:27-38 records 166,321 people, 57,240 households, 31,889 under-13 children, both native loaders, preservation of all original columns/weights, exact attendance reload, and production_ready: false. I independently hashed the seven production recipe modules; all seven hashes match the report's attendance_recipe. This closes the stale-value/relabeling path and covers the annual and final native-output boundaries. The receipt is an integrity/provenance control, not a signature or publication authorization, as documented at docs/us-childcare-attendance.md:216-230.

C2 — CRITICAL (Must Fix before merge): the required raw-source build and certification receipt have not been produced (OPEN)

  • Location: docs/us-release-build-rule.md:5-12,14-29; experiments/us-childcare-attendance/merge-readiness.md:64-108,110-134,152-154.
  • Trigger: merge a PR that materially changes how the US population is built.
  • Expected: the repository's explicit rule requires a from-scratch raw-source build and successful preflight/certification receipt from the PR's own tree before merge. Default promotion separately requires the exact-k frozen-register improvement gate; publication is a further human action.
  • Observed: the merge-readiness audit states that reusing the July BuildP parent is insufficient, the shown command is only a prepared supplied-parent integration check and was not completed, and raw-build prerequisites/adequate build resources remain outstanding. No from-scratch build or certification receipt from this head is present.
  • Impact: unit tests, standalone candidate preparation, and aggregate reports do not demonstrate that all real production inputs compose through the full builder or that required release gates pass. This is a confirmed unmet merge requirement, not merely absent optional validation. Keep the PR draft and do not merge until the prescribed raw build/certification evidence exists. This finding does not imply that the attendance mapping is wrong; it establishes that repository-defined end-to-end completion has not been demonstrated.

Should Address

A1 — SHOULD ADDRESS — RESOLVED: final release export enforces rowwise completeness and schedule validity (RESOLVED)

  • Location: packages/microcosm-build/src/microcosm/build/us_runtime/nsece_childcare.py:424-444; packages/microcosm-build/src/microcosm/build/us_runtime/childcare_attendance.py:44-58; packages/microcosm-build/src/microcosm/build/us_runtime/release_input_coverage.py:604-633; tools/build_us_fiscal_refresh_release.py:11712-11728.
  • Prior trigger: a later pipeline step could introduce a null, out-of-range value, fractional monthly count, or incoherent zero in only some rows while the generic column-nondegeneracy gate still passed.
  • Expected / observed now: the final coverage gate and unconditional pre-write fiscal boundary call assert_childcare_attendance_exportable; it requires all three columns on every row, all values finite and nonnegative, bounds of 31/7/24, integral monthly days, and coherent all-zero/nonzero schedules. They then require the content binding. Generic coverage/evidence overrides cannot waive the final check.
  • Impact: partial row loss can no longer silently fall back to engine defaults at release serialization.

A2 — SHOULD ADDRESS — RESOLVED: sibling coupling rejects missing or noncanonical household identities (RESOLVED)

  • Location: packages/microcosm-build/src/microcosm/build/us_runtime/childcare_population.py:95-119; packages/microcosm-build/src/microcosm/build/us_runtime/childcare_attendance.py:61-71,113-123,243-257.
  • Prior trigger: null or unresolved source household IDs were stringified as "nan", causing unrelated children to share one household rank.
  • Expected / observed now: household identity indices must be unique, source IDs non-null and canonical nonblank strings or non-boolean integers, and every person link must resolve before conversion to string. The imputer separately rejects blank and stringified-null values when dependence is active.
  • Impact: malformed/legacy inputs now fail rather than inducing artificial cross-household dependence. Duplicate source identities across intentional target clones remain possible, which is consistent with the clone design; the invalid/missing identity path is closed.

A3 — SHOULD ADDRESS — REFRAMED / STILL OPEN: expanded validation demonstrates that the production household schedule process fails its declared screens (STILL OPEN)

  • Location: packages/microcosm-build/src/microcosm/build/us_runtime/nsece_childcare_sibling_validation.py:139-217,219-425,588-643; packages/microcosm-build/src/microcosm/build/us_runtime/childcare_attendance_stage.py:162-174; experiments/us-childcare-attendance/review-fixes-source-stage.json:1397-1595; experiments/us-childcare-attendance/README.md:619-647; docs/us-childcare-attendance.md:111-153.
  • Trigger: the production stage fits one dependence parameter to binary attendance among the youngest complete sibling pair, then uses that shared rank to couple days and hours for all siblings, including households with three or more children.
  • Expected: household-separated evaluation of the implemented weighted donor distributions should cover participation, days, weekly hours, cross-moments/correlations, and all-child totals; a production household model should meet declared criteria or be replaced/qualified with a defensible alternative.
  • Observed: the new evaluator properly holds out whole households, integrates the actual donor CDFs, includes all children in complete 3+ child households, and marks the result non-production-ready. That resolves the prior missing diagnostic, but not the model concern. The current production model fails 9/15 provisional screens. Among 1,941 complete sibling households (661 with 3+ children), youngest-pair days correlation is 0.580 observed versus 0.442 modeled and weekly-hours correlation is 0.523 versus 0.356. For 3+ child households, mean total days are 4.510 observed versus 5.653 modeled (+25.35%) and weekly hours are 33.085 versus 40.814 (+23.36%). The code's screen thresholds are declared at nsece_childcare_sibling_validation.py:588-643; the committed failures include a 0.138 youngest-pair days-correlation gap and a 25.35% larger-household mean-days gap (review-fixes-source-stage.json:1524-1525,1587-1588).
  • Impact: child-level marginals can look plausible while nonlinear household CCDF eligibility/amounts reflect too-high larger-family intensity and too-low sibling schedule correlation. The hard household-size challenger still fails 8/15 screens; pooled matching passes 13/15 original household screens but has incompatible remaining hours moments and underpredicts hours for observed children with unresolved siblings; composition and QRF improve some means but worsen or retain the unresolved-sibling hours failure. These alternatives are explicitly diagnostic-only and are not used by childcare_attendance_stage.py:162-174. A3 is therefore statistically explored, not resolved. Its prior SHOULD ADDRESS severity is retained because these are provisional/internal screens rather than an external release gate, and the PR remains a draft candidate rather than publishing a default.

A4 — Target-frame checkpoints are not bound to the attendance receipt (OPEN)

Evidence: tools/build_us_fiscal_refresh_release.py:2200-2256 defines the target-frame identity from the original base H5 hash, engine version, seed, registry, SSI assignment, selection, and geography inputs, but has no attendance binding or execution digest. The attendance stage runs in memory before target materialization (tools/build_us_fiscal_refresh_release.py:9580-9597), while the unchanged identity is passed to the checkpoint loader at tools/build_us_fiscal_refresh_release.py:10634-10663. A checkpoint hit reconstructs stored tables solely from that identity (tools/build_us_fiscal_refresh_release.py:2321-2386).

Trigger / reproduction: run once with a durable --checkpoint-root; then change an attendance recipe module (or otherwise produce a different valid attendance binding with the same base H5, seed, PolicyEngine-US version, and other listed identity fields) and rerun against the same checkpoint root. The receipt correctly identifies the changed recipe, but _target_frame_checkpoint_identity(...) remains unchanged, so _read_target_frame_checkpoint(...) accepts the prior target frame.

Expected: any input that can alter materialized target columns must invalidate the target-frame checkpoint and reform-vector cache.

Observed: the attendance binding/execution SHA is absent from both the checkpoint identity and the cache context beginning at tools/build_us_fiscal_refresh_release.py:10665.

Impact: the calibration matrix can come from the prior attendance realization even while the exported base frame and final receipt contain the new one. The issue is immediately observable if an attendance column is named through the supported repeatable --selection-mass-protection path (tools/build_us_fiscal_refresh_release.py:950-960); it also affects any current or future materialized target/reform measure that depends on these inputs. Add the verified attendance execution/binding digest to both identities (and a regression proving the second run misses).

Suggestions

S1 — SUGGESTION — RESOLVED AS WRITTEN; transport validity remains a material evidence gap (RESOLVED)

  • Location: experiments/us-childcare-attendance/review-validation-criteria.txt:1-5; packages/microcosm-build/src/microcosm/build/us_runtime/childcare_sensitivity.py:19-126,129-181; experiments/us-childcare-attendance/review-fixes-sensitivity.json:75-98,2102-2145; docs/us-childcare-attendance.md:27-35,289-340; experiments/us-childcare-attendance/README.md:649-655.
  • Prior request: foreground the actual questionnaire-transport discrepancy, declare acceptance criteria, preserve observed regular hours, and measure benefit sensitivity to irregular-care/day assumptions.
  • Observed now: the criteria file declares the national/state flags and the documentation clearly says that days and irregular care are unidentified. The new paired evaluator requires the same donor identity, verifies the candidate equals that donor, changes only modeled bridge components, leaves measured-calendar assignments unchanged, and preserves observed regular hours. The current national changes are -3.56% without modeled irregular care, -0.04% for one fewer day, and +0.48% for one more day, below the 10% national screen; state flags remain for SD (-49.7%), TN (-36.0%), WV (-31.1%) without irregular care and WV (+38.9%) with one more day. OK's -40% flag is only a $2.06 change and is disclosed rather than interpreted as material.
  • Disposition / impact: this satisfies the original suggestion's documentation and sensitivity request, so S1 is resolved. It does not establish transport validity: the alternatives are stress tests, not confidence intervals; true days and irregular care remain unobserved, and state results remain sensitive to a few weighted donor assignments. That unresolved empirical premise is retained below as an evidence gap rather than silently converting S1 into approval.

Coordinator assessment: The code reviewer retained this as open because the transport premise itself remains unresolved. I adopt the narrower policy/source disposition: the original requested documentation, predeclared screens, and paired sensitivity analysis were supplied, so S1 is resolved as a review action; transport validity remains explicitly listed as a material evidence gap and does not become approval.

S2 — SUGGESTION — RESOLVED: the receipt records the material operations in execution order (RESOLVED)

  • Location: packages/microcosm-build/src/microcosm/build/us_runtime/childcare_attendance_stage.py:190-223; docs/us-childcare-attendance.md:200-215; committed example in experiments/us-childcare-attendance/review-fixes-population.json under candidate_receipt.childcare_attendance_stage.operations.
  • Prior trigger: three distinct source transformations were serialized as repeated generic derive_childcare_inputs labels, obscuring execution identity and order.
  • Expected / observed now: the receipt distinguishes calendar derivation, fitted sibling dependence (including rho), regular-hours schedule bridge, ASEC predictor harmonization, joint weighted schedule transfer, and outside-domain export policy in actual order. These operation records are included in the context bound by the receipt digest.
  • Impact: receipt-level auditability and future change detection are materially improved; the prior naming/organization suggestion is closed.

Evidence Gaps

  • Local affected tests were NOT RUN because the snapshot had no existing Python environment with pytest and the review contract prohibited installing dependencies; the bounded attempt and failure are recorded in pr-916-review-tests.log.
  • The licensed NSECE DS4/DS5 bytes and exact 2024 value-label codebooks were unavailable, so the detailed provider/gap/respondent/status mapping and pinned hashes could not be independently verified.
  • Noncalendar days and irregular care remain unidentified, and the inspected survey partitions are development evidence rather than independent validation; paired stress tests quantify but cannot resolve this transport premise.
  • The licensed inputs, ASEC cache, parent/candidate H5 files, and private per-person receipt inventory were unavailable, so the real candidate and 51-jurisdiction aggregate claims could not be independently reproduced; the separate required raw-source build/certification is also still outstanding as C2.

Notes

  • What looks good: the follow-up genuinely closes the prior provenance and final-boundary defects. The receipt now binds exact values to source, code, runtime and settings; native, L0 and annual paths preserve it; household IDs fail closed; and the operation sequence is auditable.
  • The PR does not add calibration targets, publish a population, or complete a full raw-source rebuild. It adds a required post-spine/fiscal-build attendance stage, release-input coverage and provenance gates, plus an attendance-only candidate rebuilt from the existing pinned parent and extensive diagnostic experiments.
  • The exact-k childcare configuration is currently optional, but the current authenticated multispine ingress does not restore the attendance receipt; omission therefore fails safely unless the supported input already arrives through a receipt-aware path. Clarifying that contract would improve the runbook but is not counted as a separate finding.
  • At the exact reviewed head the PR remains draft and GitHub reports a clean merge state.

Validation Summary

Local affected tests: NOT RUN (system Python lacks pytest; no existing project environment). Exact-head GitHub CI: 24/24 SUCCESS at 25a54a4. Reviewed the attendance-specific incremental commits and 38,155-line targeted diff; the policy role independently matched all seven committed production recipe hashes. No licensed-source or full-build reproduction was possible.

Timing

setup seconds: 77.00s; scope seconds: 33.00s; parallel review seconds: 504.00s; policy role seconds: 471.00s; code role seconds: 322.00s; adjudication seconds: 31.00s; consolidation cleanup seconds: 100.00s; elapsed seconds: 661.00s

Review Severity

REQUEST_CHANGES. Open findings: 1 critical, 2 should address, 0 suggestions.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants